Skip to content

[gemma_cpp] Add support for Qwen3 models (0.6B, 1.7B and 4B).#964

Open
copybara-service[bot] wants to merge 1 commit into
devfrom
test_952434363
Open

[gemma_cpp] Add support for Qwen3 models (0.6B, 1.7B and 4B).#964
copybara-service[bot] wants to merge 1 commit into
devfrom
test_952434363

Conversation

@copybara-service

Copy link
Copy Markdown

[gemma_cpp] Add support for Qwen3 models (0.6B, 1.7B and 4B).

Qwen3 shares a very similar architecture as Gemma3 (comparison).
Differences include:

  • Tokenizer : Qwen3 uses bytelevel BPE while Gemma3 uses Sentencepiece;
  • Embedding : Qwen3 0.6B and 1.7B do not share the input_embedding and lm_head weights, and there is no Embedding Scaling for Qwen3;
  • Activation : Qwen3 uses SiLU while gemma3 uses GeLU;
  • Norm : Qwen3 only has two pre-norms, and they are not Zero-centered (we substract 1.0 while converting the weights to keep the kernel untouched);
  • Attention : Qwen3 uses full-attention for all the layers.
  • Prompt Wrapping : Qwen3 uses different special tokens and does not require .

Qwen3 shares a very similar architecture as Gemma3 ([comparison](https://sebastianraschka.com/llm-architecture-gallery/?compare=qwen3-0-6b%2Cgemma-3-270m#architecture-diff-tool)).
Differences include:
- *Tokenizer* : Qwen3 uses bytelevel BPE while Gemma3 uses `Sentencepiece`;
- *Embedding* : Qwen3 0.6B and 1.7B do not share the `input_embedding` and `lm_head` weights, and there is no Embedding Scaling for Qwen3;
- *Activation* : Qwen3 uses SiLU while gemma3 uses GeLU;
- *Norm* : Qwen3 only has two pre-norms, and they are not Zero-centered (we substract 1.0 while converting the weights to keep the kernel untouched);
- *Attention* : Qwen3 uses full-attention for all the layers.
- *Prompt Wrapping* : Qwen3 uses different special tokens and does not require <BOS>.

PiperOrigin-RevId: 952434363
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants